Papers with multi-scale visual information

2 papers
Multi-Scale Progressive Attention Network for Video Question Answering (2021.acl-short)

Copied to clipboard

Challenge: Experimental evaluations on three benchmarks: TGIF-QA, MSVD-QA and MSRVTT-QA show our method has achieved state-of-the-art performance.
Approach: They propose a multi-scale progressive attention network to fuse visual and text information.
Outcome: The proposed method achieves state-of-the-art on three benchmarks: TGIF-QA, MSVD-QA and MSRVTT-QA.
Multimodal Sarcasm Target Identification in Tweets (2022.acl-long)

Copied to clipboard

Challenge: Existing methods to detect sarcasm target with text lacking context are not sufficient and complete.
Approach: They propose a multi-modal sarcasm target identification task that performs both textual and visual detection.
Outcome: The proposed model can perform textual target labeling and visual target detection.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations